Papers with conditional graph attention network
Are Scene Graphs Good Enough to Improve Image Captioning? (2020.aacl-main)
Copied to clipboard
| Challenge: | Existing image captioning models rely on object detection features to generate image descriptions, but they are noisy. |
| Approach: | They propose to use scene graphs to introduce information about object relations into captioning to improve image descriptions. |
| Outcome: | The proposed model improves image caption quality by 3.3 CIDEr compared to a strong Bottom-Up Top-Down baseline. |